keywords:"bilingual document alignment" - Search Results - Digital Repository

guest :: login Digital Repository
		Search		Submit		Help		About

Home > Search Results: keywords:"bilingual document alignment"

Search:

Search Tips :: Advanced Search

Search collections:

Sort by:	Display results:	Output format:

	Mining Parallel Corpora from the Web Kúdela, Jakub ; Holubová, Irena (advisor) Title: Mining Parallel Corpora from the Web Author: Bc. Jakub Kúdela Author's e-mail address: jakub.kudela@gmail.com Department: Department of Software Engineering Thesis supervisor: Doc. RNDr. Irena Holubová, Ph.D. Supervisor's e-mail address: holubova@ksi.mff.cuni.cz Thesis consultant: RNDr. Ondřej Bojar, Ph.D. Consultant's e-mail adress: bojar@ufal.mff.cuni.cz Abstract: Statistical machine translation (SMT) is one of the most popular ap- proaches to machine translation today. It uses statistical models whose parame- ters are derived from the analysis of a parallel corpus required for the training. The existence of a parallel corpus is the most important prerequisite for building an effective SMT system. Various properties of the corpus, such as its volume and quality, highly affect the results of the translation. The web can be considered as an ever-growing source of considerable amounts of parallel data to be mined and included in the training process, thus increasing the effectiveness of SMT systems. The first part of this thesis summarizes some of the popular methods for acquiring parallel corpora from the web. Most of these methods search for pairs of parallel web pages by looking for the similarity of their structures. How- ever, we believe there still exists a non-negligible amount of parallel... Detailed record
	Mining Parallel Corpora from the Web Kúdela, Jakub ; Holubová, Irena (advisor) ; Helcl, Jindřich (referee) Title: Mining Parallel Corpora from the Web Author: Bc. Jakub Kúdela Author's e-mail address: jakub.kudela@gmail.com Department: Department of Software Engineering Thesis supervisor: Doc. RNDr. Irena Holubová, Ph.D. Supervisor's e-mail address: holubova@ksi.mff.cuni.cz Thesis consultant: RNDr. Ondřej Bojar, Ph.D. Consultant's e-mail adress: bojar@ufal.mff.cuni.cz Abstract: Statistical machine translation (SMT) is one of the most popular ap- proaches to machine translation today. It uses statistical models whose parame- ters are derived from the analysis of a parallel corpus required for the training. The existence of a parallel corpus is the most important prerequisite for building an effective SMT system. Various properties of the corpus, such as its volume and quality, highly affect the results of the translation. The web can be considered as an ever-growing source of considerable amounts of parallel data to be mined and included in the training process, thus increasing the effectiveness of SMT systems. The first part of this thesis summarizes some of the popular methods for acquiring parallel corpora from the web. Most of these methods search for pairs of parallel web pages by looking for the similarity of their structures. How- ever, we believe there still exists a non-negligible amount of parallel... Detailed record
	Mining Parallel Corpora from the Web Kúdela, Jakub ; Holubová, Irena (advisor) Title: Mining Parallel Corpora from the Web Author: Bc. Jakub Kúdela Author's e-mail address: jakub.kudela@gmail.com Department: Department of Software Engineering Thesis supervisor: Doc. RNDr. Irena Holubová, Ph.D. Supervisor's e-mail address: holubova@ksi.mff.cuni.cz Thesis consultant: RNDr. Ondřej Bojar, Ph.D. Consultant's e-mail adress: bojar@ufal.mff.cuni.cz Abstract: Statistical machine translation (SMT) is one of the most popular ap- proaches to machine translation today. It uses statistical models whose parame- ters are derived from the analysis of a parallel corpus required for the training. The existence of a parallel corpus is the most important prerequisite for building an effective SMT system. Various properties of the corpus, such as its volume and quality, highly affect the results of the translation. The web can be considered as an ever-growing source of considerable amounts of parallel data to be mined and included in the training process, thus increasing the effectiveness of SMT systems. The first part of this thesis summarizes some of the popular methods for acquiring parallel corpora from the web. Most of these methods search for pairs of parallel web pages by looking for the similarity of their structures. How- ever, we believe there still exists a non-negligible amount of parallel... Detailed record

Interested in being notified about new results for this query?
Subscribe to the RSS feed.

Digital Repository :: :: :: ::
Powered by v1.1.2
Maintained by

This site is also available in the following languages:
Česky English